Papers with visually grounded tasks

2 papers
LaMI: Augmenting Large Language Models via Late Multi-Image Fusion (2026.acl-short)

Copied to clipboard

Challenge: Large Language Models lack visual grounding on visual reasoning, despite training on text alone.
Approach: They propose a late multi-image fusion method that augments LLMs with test-time visual signals.
Outcome: Using a late multi-image fusion method, the proposed model outperforms LLMs on visual reasoning and matches VLMs in vision-based tasks.
Symmetrical Visual Contrastive Optimization: Aligning Vision-Language Models with Minimal Contrastive Images (2025.acl-long)

Copied to clipboard

Challenge: Recent studies have shown that Large Vision-Language Models (VLMs) tend to neglect image content and over-rely on language-model priors, resulting in errors in visually grounded tasks and hallucinations.
Approach: They propose a novel finetuning objective that steers the model toward capturing important visual details and aligning them with corresponding text tokens.
Outcome: The proposed method achieves up to 22% reduction in hallucinations and significant gains in vision-centric and general tasks while maintaining or improving the model's general abilities.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations